Job Title: Lead, DevOps and Platform Engineering
Company: Kotak Securities Limited
Location: Bangalore
Experience: 10–15 years
About the Role
Lead the platform engineering function supporting Kotak Securities’ technology transformation across trading, risk, market data, observability, data platforms, back-office systems, and datacenter migration. Build the platform team and own compute infrastructure, delivery pipelines, infrastructure definitions, and instrumentation standards.
Key Responsibilities
• Define and manage compute platforms across Kubernetes and virtual machines, including cluster architecture, networking, workload isolation, and resource governance.
• Own infrastructure as code across cloud and on-premise environments, including reusable modules, state management, drift detection, and policy enforcement.
• Establish performance standards for latency-sensitive trading workloads, including CPU pinning, NUMA awareness, kernel tuning, and network optimization.
• Secure the container supply chain through image hardening, vulnerability checks, artifact signing, and provenance.
• Design CI/CD and GitOps pipelines with progressive delivery, trading-aware deployment windows, and rollback strategies.
• Establish service level objectives, error budgets, incident management, and postmortem practices integrated with change management.
• Own observability architecture, including instrumentation, tracing, metrics, logging, retention, and ITSM integration.
• Drive capacity planning and performance testing for peak market events and pre-open traffic surges.
• Lead platform workstreams for datacenter migration, including hybrid connectivity, environment parity, and cutover planning.
• Own platform security, secrets management, certificates, workload identity, privileged access, and audit readiness.
• Optimize infrastructure costs through rightsizing, commitment planning, and service-level cost attribution.
• Build, lead, and mentor the platform engineering team.
Technical Skills
• Deep production expertise in Kubernetes, including control plane operations, upgrades, CNI, service mesh, scheduling, autoscaling, and failure analysis.
• Strong experience with Terraform or equivalent infrastructure-as-code tools at scale.
• Strong cloud networking and security expertise on AWS or comparable platforms, including VPCs, private connectivity, load balancing, DNS, and IAM.
• Deep Linux systems knowledge, including kernel tuning, TCP behavior, filesystems, I/O, and diagnostics using perf or eBPF-based tools.
• Production programming experience in Go, Python, or an equivalent language for platform tooling, operators, or controllers.
• Experience designing organization-wide CI/CD and GitOps workflows using ArgoCD, Flux, or equivalent tools.
• Strong understanding of OpenTelemetry, APM, logging pipelines, trace sampling, and metric cardinality management.
• Experience with policy-as-code tools such as OPA or Sentinel.
Required Experience
• 10 or more years in infrastructure, platform engineering, or site reliability roles, including at least 3 years leading engineers.
• Experience operating critical systems where downtime has direct commercial or regulatory consequences.
• Proven ability to deliver within audit, change control, and segregation of duties requirements.
Preferred Experience
• Capital markets or payments infrastructure, including exchange connectivity, market data distribution, colocation, and clock synchronization.
• Large-scale datacenter or colocation migrations, including hybrid routing, BGP, and firewall transitions.
• Production ownership of Postgres, Redis, Kafka, or MongoDB, including replication, failover, and backup verification.
• Exposure to SEBI CSCRF, RBI cybersecurity guidelines, or comparable frameworks.
• Experience building an observability practice from the ground up.
Why Kotak Securities
• Business-funded technology transformation with a multi-year roadmap.
• Full-time opportunities on a long-term, in-house engineering team.
• Internal mobility across banking, broking, asset management, and insurance.
• Structured recognition, long-service, and returnship programmes.
• An engineering environment focused on auditability, resilience, and regulatory governance.
• A culture grounded in approachability, mutual respect, transparency, entrepreneurial thinking, and ethical conduct.